Genome Research
● Cold Spring Harbor Laboratory
Preprints posted in the last 7 days, ranked by how well they match Genome Research's content profile, based on 468 papers previously published here. The average preprint has a 0.27% match score for this journal, so anything above that is already an above-average fit.
Patsakis, M.; Tzanakakis, A.; Georgakopoulos-Soares, I.
Show abstract
Evo 2 is the largest openly available genomic foundation model, but its forty billion parameter configuration cannot be loaded onto a single 80 GB accelerator, placing genome-scale analysis beyond most laboratories. We present TurboQuant-Bio, an open toolkit that compresses Evo 2s weights and attention cache to four bits without calibration data, and serves both through fused kernels. Compression is near-lossless across perplexity spanning the tree of life, genomic classification, splice-site prediction, gene completion and clinically relevant variant-effect prediction. It brings Evo 2 40B onto one 80 GB GPU and Evo 2 7B to its full million-token context within a 40 GB memory budget, an eightfold gain in reachable context. We further show that the released chunked-prefill path is silently incorrect, returning plausible but uncorrelated likelihoods, and derive the block-wise continuation that repairs it: a complete 580-kilobase bacterial genome is now scored in one context in 22 minutes rather than 13.7 hours.
Fu, Y.; Morley, C.; Masters, L. M.; English, A. C.; Zhu, Y.; Moller, A. G.; Paulin, L. F.; Thompson, B.; Kalef-Ezra, E.; Weissenberger, G.; Shen, H.; Meridith, M.; Manini, A.; Horner, D.; Reed, X.; Muzny, D.; Jaunmuktane, Z.; Khan, Z. M.; Mehta, H.; Timp, W.; Billingsley, K.; Erwin, G. S.; Proukakis, C.; Sedlazeck, F. J.
Show abstract
Somatic mutations arise throughout life, with functional consequences tied to the cell populations in which they occur. Genome-wide studies measure somatic variations in bulk tissue, whereas single-cell approaches resolve cell identity but provide limited sensitivity for complex alleles. Here we developed SniffCell, which uses DNA methylation carried on native long reads to assign somatic variant-supporting molecules to methylation-resolvable cell types. SniffCell builds cell-type-discriminatory methylation signatures across eight tissues, assigns long reads to cell types, and provides cell-type-specific variant calling. Across peripheral blood mononuclear cells and brain benchmarks, SniffCell recovered sorted cell identities and validated cell-type-specific variant assignments using purified immune-cell, neuronal, and oligodendrocyte fractions. In blood, SniffCell recovered lineage-restricted antigen receptor rearrangements and localized a somatic tandem-repeat expansion to T cells. In the frontal cortex, SniffCell identified recurrent neuron-specific tandem-repeat expansions in genes including FGF14, LRRC7 and SH3RF3. Across three brain cohorts comprising 172 donors, recurrent neuron-associated expansions were enriched for GAA-rich motifs. In donors with matched blood, and diverged more strongly from the inherited repeat length, whereas oligodendrocyte-associated alleles more often tracked it. SniffCell transforms native bulk long-read genomes into a cell-type-aware resource for somatic variant discovery and reveals recurrent somatic instability in human tissues at cell-type resolution.
Wiel, L.; Ferraro, F.; Yu, J.; Zhen, J.; Nachun, D.; Mendez, R.; Reuter, C. M.; Cui, J. L.; Bonner, D. E.; Carter, J. N.; Marwaha, S.; van de Vorst, M.; Emami, S.; Kravets, E.; Neu, M. B.; van Ham, T. W.; Kleefstra, T.; Ashley, E. A.; Bernstein, J. A.; Montgomery, S. B.; Gilissen, C.; Wheeler, M. T.
Show abstract
The interpretation of missense variants remains a major challenge in clinical genetics. "Meta-domains" aggregate population and pathogenic variation across homologous Pfam domain instances in the human proteome, providing per-residue context for interpreting variants of uncertain significance (VUS). Our 2019 implementation, MetaDome, is widely used and named in clinical variant-classification guidelines. Here we present the MetaDome 2027 update, featuring a comprehensively updated dataset and GRCh38 support. The redesigned pipeline enables incremental updates of GENCODE, UniProtKB/Swiss-Prot, Pfam, gnomAD, and ClinVar while maintaining 100% sequence-identity gene-to-protein mapping. Annotated Pfam domain instances grew 14.9% from 71,419 to 82,069 and meta-domain-eligible Pfam families ([≥]2 human occurrences) by 73.3% from 3,334 to 5,778; Pfam domains are annotated to 92% of human proteins. Approximately 43% of mapped protein-coding nucleotides (14.3 million in GRCh38, 13.8 million in GRCh37) are in a meta-domain; in GRCh38 67.9% (37,692 of 55,548) of pathogenic or likely pathogenic ClinVar missense variants fall at such a position. We show how MetaDome helped reclassify a de novo missense VUS in RALA and identify 52,463 ClinVar missense VUS for which meta-domains supply otherwise unavailable pathogenic evidence. MetaDome is freely available at www.metadome.app.
Ding, Y.; Zhang, P.; Ociepa, T.; Nucia, A.; Guan, H.; Kowalczyk, K.; Park, R. F.; Okon, S.
Show abstract
Blumeria graminis f. sp. avenae (Bga), the causal agent of oat powdery mildew, is one of the most host-specialized members of the B. graminis species complex. Despite its agricultural importance, the lack of a high-quality reference genome has limited studies of host specialization, virulence evolution and comparative genomics in this pathogen. Here, we generated the first chromosome-scale genome assembly of Bga using an integrative approach combining long- and short-read sequencing, Hi-C scaffolding and transcriptome data. The Bga genome exhibits hallmark features of powdery mildew fungi, including extensive repeat content and low gene density. Comparative analyses revealed that genome expansion is primarily associated with historical transposable element proliferation rather than recent transpositional activity. Genome organization is consistent with a functionally stratified "one-speed" model, in which genes associated with pathogenicity, including predicted effectors and infection-responsive genes, are preferentially located in transposable element-rich regions characterized by reduced synteny conservation and extended intergenic spaces. In contrast, conserved genes are concentrated in compact genomic regions and maintain strong syntenic conservation across cereal-infecting formae speciales. Hi-C analyses demonstrated a highly structured chromatin architecture and revealed genome organization patterns associated with infection-related gene expression. Comparative genomic analyses indicated that host specialization in Bga is driven by localized diversification of a relatively small subset of genes rather than large-scale genome restructuring. These results provide the first high-quality genomic resource for Bga and offer new insights into the evolutionary mechanisms underlying host specialization in powdery mildew fungi.
Dutta, S.
Show abstract
Bulk telomere length measured from tumour sequencing is routinely interpreted as a property of the cancer cells. However, a tumour specimen is a mixture, and the patient who supplies it has a telomere length of their own. Here I re-analyse published pan-cancer telomere estimates and ask how much of a tumour's telomere length is patient-specific. A calibration step comes first. Whole-genome and low-pass estimates recover the known cross-sectional attrition of leukocyte telomeres with age, at 26.6 bp per year in blood normals, whereas whole-exome estimates do not. After adjustment for cancer type, sequencing centre and sex, the exome slope is minus 0.6 bp per year. In 684 blood-normal aliquots sequenced by both assays, the whole-genome estimate declines at 38.9 bp per year, whereas the exome estimate from the same DNA shows no detectable decline. The difference between assays is 41.5 bp per year, with P = 3 x 10^-10. Because exome data constitute 78.6% of the original resource, downstream analyses use only whole-genome and low-pass libraries. Within those data, tumour telomere length tracks the patient's matched-normal telomere length. The Spearman correlation is 0.395 in TCGA, with positive associations in 22 of 23 cancer types. This finding replicates in PCAWG using a different telomere estimator, with a correlation of 0.472 and positive associations in all 24 histologies examined. Adjustment for cancer type, sequencing centre and library type leaves a regression coefficient of 0.385. The association is also stable after adjustment for age, sex, tumour purity, leukocyte fraction, ploidy, sequencing coverage and continental ancestry, with coefficients ranging from 0.406 to 0.429. Pure normal-cell admixture is rejected as the sole explanation. Under a two-compartment mixture model, the coefficient for host telomere length is expected to equal 1 and the host-by-purity interaction to equal minus 1. These restrictions are jointly rejected with P = 0.001. Tumour purity, leukocyte fraction and age each explain only about 1 to 3% of within-cohort variance and do not alter the cross-cancer ranking. By contrast, the between-cohort coefficient is not directly interpretable. Its apparent near one-to-one relationship with tissue-associated telomere length depends strongly on which tissue supplies the matched-normal reference and on the statistical spread of that predictor, falling to 0.44 when organ-matched solid tissue is used. Bulk tumour telomere length is therefore a composite phenotype containing a replicated patient-specific component. Telomere biomarker studies should include matched-normal telomere length as a covariate rather than treating tumour telomere length as exclusively tumour-intrinsic.
Su, X.; Peng, Y.; Yang, X.; Zhang, F.; Xu, Q.; Ma, Z.; Dong, Y.; Zhou, L.; Xue, H.; Cao, X.; Zou, Z.; Wang, Y.; Zhou, Y.; Zeng, X.
Show abstract
Oil palm (Elaeis) is the primary source of global vegetable oil. Interspecific hybrids of Elaeis exhibit pronounced heterosis by integrating two distinct subgenomes into a single nucleus, effectively combining the high yield of African oil palm (E. guineensis) with the high unsaturated fatty acid content and disease resistance of American oil palm (E. oleifera). However, the genetic basis underlying heterosis is still unclear. Here, we combine phased genome assembly, comparative genomics, evolutionary genomics and haplotype-aware transcriptomics to unravel the genetic architecture of heterosis of hybrid oil palm. We assemble the highly heterozygous F1 genome ('Reyou 40', 3.75% heterozygosity) into a complete 1.73 Gb T2T haplotype (HapG) and a 1.84 Gb near-T2T haplotype (HapO with17 gaps). Despite 91.56% sequence identity, HapG and HapO diverged in LTR-RT occurrence and PAV affected genes, showing complementary biases in lipid metabolism and stress responses, respectively. Evolutionary genomics revealed that ancient WGDs preserved the palm family. Whereas lineage-specific lipid-related gene expansions in oil palm. Six ancient introgressed regions (~64 Mb) in HapG were reshaped by transposable elements and tandem duplication, showing an enrichment of genes related to resistance and lipid metabolism. Transcriptomically, 82.2% of allelic gene pairs maintained balanced expression, accompanied by parental functional complementarity and dosage buffering, revealing a potential regulatory basis for coordinating parental genetic differences in the hybrid genome. These haplotype-resolved genomic resources offer vital targets for understanding heterosis and accelerating oil palm molecular breeding.
Barron, W. C.; Wei, X.; Ferdousy, S.; Zhu, L.; Meng, F. W.; Chen, B.
Show abstract
Pre-mRNA splicing is essential for gene expression, yet how disruption of core spliceosomal factors produces tissue- and developmental stage-specific phenotypes remains poorly understood. Here, we investigated the in vivo function of the conserved spliceosomal kinase PRPF-4 in C. elegans using endogenous reporter analysis, conditional protein depletion, and transcriptome-wide analysis of alternative splicing and gene expression. We found that PRPF-4 is broadly expressed throughout development and is continuously required for postembryonic development, with distinct requirements in the pharynx, nervous system, and germline. Acute PRPF-4 depletion rapidly disrupts alternative splicing across thousands of transcripts, with exon skipping representing the predominant class of affected events. In addition, PRPF-4 depletion results in a robust transcriptome shift with induction of components of the spliceosome and repression of ciliary and ion transport-related transcripts. These findings establish PRPF-4 as a central regulator of RNA metabolism and demonstrate the far-reaching effects on gene expression caused by loss of core spliceosomal components.
Li, X.; Wei, P.
Show abstract
Causal mediation analysis is widely used to identify biological pathways linking exposures to outcomes, but most methods assume homogeneous mediation effects across individuals. In high-dimensional omics settings, this assumption can mask important heterogeneity driven by demographic, genetic, or environmental factors. We propose the M-high-learner, a flexible framework for detecting heterogeneous mediation effects with high-dimensional mediators. The method identifies mediators with subgroup-specific indirect effects while distinguishing them from null or homogeneous signals and controlling the type I error rate. It is computationally efficient, scalable, and yields interpretable sub-types. Simulation studies show that the proposed approach achieves high power while maintaining accurate error control. Applications to the Framingham Heart Study and the Multi-Ethnic Study of Atherosclerosis reveal that the mediation role of gene expression in sexs effect on high-density lipoprotein varies across subgroups defined by body mass index and age. Our framework provides a practical tool for uncovering heterogeneous biological mechanisms in high-dimensional genomic studies. Author SummaryBiological processes linking risk factors to disease often differ across individuals, but many existing methods assume these processes are the same for everyone. This can hide important differences between groups. We developed a powerful method to identify when these pathways vary across subgroups using large-scale molecular data. Our approach detects differences in how intermediate biological factors contribute to outcomes in populations defined by characteristics such as age and body mass index. Applying our method to population studies, we found that some biological pathways operate differently across groups, suggesting that key mechanisms may be missed when differences are ignored. Our work provides a tool to better understand how disease-related processes vary across individuals, which may support more targeted and personalized approaches to health research.
Eliscu, R.; Kang, G.; Schupp, P. G.; Brody, D. J.; Hariharan, N.; Shamsian, S.; Oldham, M. C.
Show abstract
Genome-wide coexpression analysis of intact tissue samples is a powerful approach for identifying reproducible signatures of cell types and states, since it can survey vast numbers of individuals, cells, and transcripts. However, it can be difficult to optimize gene coexpression network construction and compare results from independent analyses. To address these challenges, we developed OMICON (theomicon.ucsf.edu) for research on human brain gene coexpression networks. OMICON contains gene expression data from >17K normal and neoplastic human brain samples with standardized metadata. Systematic analysis of independent datasets identified >250K gene coexpression modules, which were characterized and compared via enrichment analysis with >40K gene sets. All modules are discoverable via an advanced search engine that can filter by genes, metadata, and enrichment results. Analyses can also be browsed with an interactive workflow visualization tool, and users can communicate within OMICON using @mention functionality to support communal research on human brain gene coexpression networks.
Kang, G.; Oldham, M. C.
Show abstract
Understanding which genes are reproducibly dysregulated in which cell types is foundational knowledge for efforts to slow or reverse pathologies. For neuropathologies, such efforts rely primarily on differential expression analysis of single-nucleus RNA-seq (snRNA-seq) data. However, this strategy suffers from experimental and statistical challenges that limit marker gene reproducibility. We describe a novel strategy called Covariation Projection Analysis (CoPA) that combines the power of bulk sampling with the precision of single-cell methods. By projecting bulk gene coexpression modules onto pseudobulked snRNA-seq cell types, CoPA reveals the cellular origins of highly reproducible genomic programs and their relative importance among cell types. By comparing CoPA projection patterns between normal and pathological human brain samples using differential CoPA (dCoPA), we identify gene coexpression modules that are uniformly and reproducibly dysregulated in specific neocortical cell types in Alzheimers disease or schizophrenia. We share our findings through a novel web application called CoPA Cabana (https://oldhamlab.shinyapps.io/copacabana/).
Zhao, L.; Zeng, Y.; Abelman, D. D.; Lin, W.; Luo, P.
Show abstract
Motivation: Cell-free DNA methylation provides a minimally invasive signal for early cancer detection and tissue-of-origin prediction. Most methods represent methylation measurements as independent fixed-window features and therefore do not explicitly model relationships among genomic regions. Results: We developed PANGEM (Pan-cancer Graph-based Cancer Detection Using the Cell-free DNA Methylome), a graph-learning framework that represents genomic bins as nodes and integrates CpG context, genomic proximity, and sample-specific methylation similarity in the graph topology. Across five repeated stratified train-test splits, PANGEM achieved the highest mean performance among evaluated methods, with an AUROC/AUPR of 0.997/1.000 for binary cancer detection and macro-AUROC/AUPR of 0.977/0.870 for multiclass tissue-of-origin prediction. In the independent INSPIRE cohort, 72 of 78 cancer cases (92.3%) exceeded the binary classification threshold, and PANGEM correctly classified 9 of 17 head and neck cancer cases (52.9%), the highest accuracy among evaluated methods. Subnetwork analysis further identified recurrent, graph-connected methylation patterns, including a 111-DMR subnetwork with increased methylation in cancer samples.
Ivankovic, F.; Ko, A.; Aster, M. M.; Balaconis, M. K.; Banks, E.; Bemis, M.; Cibulskis, K. R.; Degatano, K.; Gauthier, L. D.; Grant, G.; Hatcher, A.; Kachulis, C.; Karczewski, K. J.; Labrecque, S. M.; Lawson, J.; Liao, C.; Magner, R.; Munshi, R.; Schatz, M. C.; Schultz, P. M.; Shah, S. P.; Sheets, E. A.; Tibbetts, K.; Vernest, K. A.; Ye, R.; Gabriel, S.; Lennon, N. J.; Neale, B. M.; Browning, B. L.; Lichtenstein, L. T.
Show abstract
Genotype imputation remains essential for large-scale human genetics studies, but its performance is limited by the size and ancestral diversity of available reference panels, reducing accuracy for rare variants and underrepresented populations. Here, we present a cloud-based imputation service built on a multi-ancestry reference panel derived from 515,579 jointly phased genomes from the All of Us (N=414,830) and National Human Genome Research Institute's Analysis, Visualization, and Informatics Lab-space (AnVIL, N=100,749) datasets. The All of Us + AnVIL reference panel is highly diverse and includes 261,163 participants most genetically similar to non-European reference populations, spanning 665,398,839 high-quality autosomal sites, representing a nearly 50% increase over TOPMed, the previous largest imputation service. Across multiple ancestry groups, the panel enables accurate imputation (empirical R2 0.8) for variants with allele frequencies as low as 0.2%, extending reliable imputation into the rare-variant frequency spectrum, including allele frequencies down to 0.002% and 0.006% for samples with European ancestry and African ancestry in the United States, respectively. Compared with TOPMed, the panel improves imputation accuracy across all ancestry groups except Africans, and recovers additional trait-associated variants not represented in existing reference panels. To facilitate broad community access while preserving participant privacy, we deploy the panel through a secure cloud-based imputation platform using privacy-preserving recombined haplotypes. This resource establishes a new foundation for genome-wide association studies (GWAS) and fine-mapping, especially in previously underrepresented populations.
Wapenaar, H.; Clifford, G.; Taglini, F. T.; McGhie, F.; Rolls, W.; Zhang, Y.; Sproul, D.; Wilson, M. D.
Show abstract
DNMT3A is a de novo DNA methyltransferase whose recruitment to chromatin regulates its function. Missense mutations within the chromatin-binding PWWP domain are associated with diverse human disorders, yet how mutations in the same domain produce distinct phenotypes remains unclear. Here we systematically characterise 19 clinically reported mutations in the PWWP domain of DNMT3A that are associated with Heyn-Sproul-Jackson syndrome (HESJAS), paraganglioma (PG) and clonal haematopoiesis (CH). We show that all PWWP-domain mutations associated with HESJAS abolished interaction with H3K36me2 modified nucleosomes, defining this as a consistent biochemical feature of HESJAS. In contrast, mutations from all disease classes differentially altered DNA binding of the PWWP domain, driven by alterations in the net charge of the domain. However, these effects are largely overcome by inclusion of the DNNMT3A1 N-terminal region, which is absent from its embryonic isoform, suggesting that PWWP mutations may differentially affect DNMT3A function through development. Changes in the thermal stability of the isolated PWWP domain mutants did not directly translate into altered stability of full-length DNMT3A1 in cells. We show that HESJAS mutations can affect the intramolecular interaction between the PWWP and adjacent ADD domain, an interaction proposed to contribute to the autoinhibitory function of the ADD domain. However, not all mutations behaved in the same way, suggesting that multiple factors govern the intramolecular autoinhibition of DNMT3A. Together, this study advances our understanding of the molecular mechanisms by which DNMT3A PWWP-domain mutations are mechanistically heterogeneous, providing a biochemical framework that contributes to distinct disease phenotypes.
Seiler, E.; Willemsen, M.; Piro, V. C.; Reinert, K.
Show abstract
Motivation: A continued decrease in sequencing costs has facilitated the exponential increase in available sequencing data, with public databases like the European Nucleotide Archive (ENA) and Sequence Read Archive (SRA) reaching well in the order of petabases. This has been the incentive to develop more scalable tools for common bioinformatics tasks. One such task is the approximate searching of short sequence patterns like genes or reads in reference data sets. In recent years, a variety of indexing data structures have been proposed for searching large sequencing databases. The state-of-the-art index, the Hierarchical Interleaved Bloom Filter (HIBF) was first-in-class to index one million samples. To be useful for expanding repositories, it must be extended to support dynamic updates. Results: In this paper, we introduce a scalable and updatable sequence-search index by extending the HIBF with partial rebuilding to support efficient updates. We demonstrate the Dynamic HIBF's capacity for large-scale data by iteratively creating an index from over 100 TB of compressed reads across more than 39,000 full human RNA-Seq samples, updated in consecutive batches of 100. To benchmark against state-of-the-art tools, we evaluated incremental performance on a subset of 5,000 samples sub-sampled to 1% of their original read depth. In this comparative setting, the dynamic HIBF completed the sequential insertion of all 5,000 samples within 5 hours--24 to 65 times faster than competing methods and twice as fast as the static HIBF.
Helekal, D.; Blomqvist, S. O. P.; Mukherjee, A.; Bowcutt, B. A.; Palace, S. G.; Grad, Y. H.
Show abstract
Bacterial genome-wide association studies (GWAS) offer a powerful approach to identify the genetic basis of a trait measured in a set of sequenced isolates. As the number of sequenced isolates has grown, the limiting factor for GWAS has become phenotyping enough isolates to achieve statistical power. To overcome the need for large-scale phenotyping, we developed Bayesian Adaptive Sequential Sampling GWAS (BASS-GWAS), which couples Bayesian adaptive experimental design with a sparse regression model to select maximally informative isolates for phenotypic testing. BASS-GWAS efficiently recovered causal loci for three antimicrobial resistance traits in Neisseria gonorrhoeae, requiring many fewer phenotyped isolates than random sampling. We applied BASS-GWAS to discover variants enabling gyrBD429N-dependent cross-resistance to the novel topoisomerase inhibitors zoliflodacin and gepotidacin. After phenotyping fewer than 30 isolates, we identified and then validated both parCD86N and a gyrA-parE-based pathway as enabling cross-resistance. BASS-GWAS provides a practical and statistically principled solution for efficient bacterial GWAS.
Tassios, E.; Pyrgelis, N.; Rinker, D.; Tzermpou, E. M.; Hittinger, C. T.; Rokas, A.; Nikolaou, C.; Vakirlis, N.
Show abstract
Genes encoding novel protein sequences are a ubiquitous feature of genomes. They fuel molecular and cellular evolutionary innovations and frequently contribute to species-specific characteristics. We are now unravelling the processes by which they originate, including de novo from noncoding sequences and through extreme divergence, yet how much and what types of novel proteins evolve through each process is still unclear Does the mechanism of origination shape the structural and functional potential of the resulting proteins? Here, we conducted a broad computational investigation of genetic and protein novelty at the scale of the entire subphylum of Saccharomycotina yeasts. We detected more than 5,000 robust de novo genes across 332 species and compared them to more than 10,000 novel genes resulting from extreme sequence divergence, revealing two distinct modes of evolution of novelty. A remarkable 40% of de novo proteins are predicted to localize to mitochondria compared to only 15% of divergent, with the latter also being substantially longer and more disordered. A detailed analysis of conservatively predicted tertiary structures of novel proteins shows that "invention" of novel folds can happen through both processes but is more likely to occur de novo. We also illustrate cases of evolutionary "re-invention" of existing protein folds from non-coding sequences. Our work deepens our understanding of the origins and importance of novel proteins opening new directions for further structural and functional characterization.
Vigna, A.; Harrouard, J.; Miot-Sertier, C.; Loegler, V.; Marullo, P.; Friedrich, A.; Schacherer, J.; Peltier, E.; Albertin, W.
Show abstract
Brettanomyces bruxellensis is a yeast species associated with diverse fermentation environments and characterized by extensive genetic diversity, including diploid, autotriploid, and allotriploid lineages resulting from independent hybridization events. These lineages are associated with distinct ecological niches and provide a framework for studying metabolic trait evolution in complex genomes. Nitrate assimilation is a relatively uncommon trait among yeasts and has been reported in B. bruxellensis, but its distribution and evolutionary history within the species remain poorly understood. Here, we combined phenotypic characterization of 151 strains with genomic analyses of 946 whole-genome sequences to investigate nitrate assimilation. Growth assays revealed that nitrate assimilation is widespread but unevenly distributed across genetic lineages, with some populations largely retaining the trait whereas others have frequently lost it. Genomic analyses identified extensive variation affecting the nitrate assimilation gene cluster composed of YNR1, YNI1, and YNT1. Nitrate assimilation was strongly associated with both gene copy number and predicted gene functionality, with nitrate-assimilating strains generally carrying more functional copies of the cluster. Leveraging the complex genomic architecture of the species, we independently analyzed primary and acquired genomes in allotriploid lineages and uncovered contrasting evolutionary trajectories following hybridization. While nitrate assimilation genes were generally maintained in primary genomes, acquired genomes showed a higher prevalence of gene loss and predicted loss-of-function variants, revealing asymmetric dynamics between subgenomes. Altogether, our results suggest that nitrate assimilation represents an ancestral trait that has been differentially maintained across B. bruxellensis lineages through a combination of copy number variation, gene degeneration, and genome-specific evolutionary dynamics. These findings provide new insights into how genome architecture and polyploid evolution shape the maintenance and loss of metabolic traits in an industrially relevant yeast species.
Cabanas, N.; Veloso, A.; Zinzen, R.; Bucher, G.
Show abstract
The brain is essential for animal survival and based on its conserved Bauplan, an impressive adaptive diversity has evolved. However, the genetic mechanisms regulating brain development and diversification remain enigmatic. The insect neural stem cells (neuroblasts, NBs) acquire different identities through the combinatorial expression of transcription factors (TFs), but this code is unknown for the brain. Here, we define the conserved core of TFs expressed in insect brain NBs by a combined analysis of single-cell expression from NBs derived from two holometabolous insects, the fly Drosophila melanogaster and the beetle Tribolium castaneum. In Tribolium, we established a Gal4 enhancer trap system to identify a line that marks NBs. From 37,137 sequenced NBs, we identified 10,425 brain NBs. In Drosophila, we sequenced 32,112 NB nuclei, identifying 12,389 brain NBs. Analysing the combined dataset strongly increased the sensitivity in specifying the core of 188 brain-specific TFs. We found two atypical clusters with some similarity to Type II NBs and identified seven transcription factors not previously associated with or confirmed in NBs (Hmx, CG15696, CG32532, dmrt99B, fD59A, TfAP-2, and Fer1). Our data reveals fundamental differences between brain and ventral nerve cord specification and paves the way to study the development and evolution of brain specific structures.
Velazquez, D.; Hallinan, C.; An, R.; Clifton, K.; Fan, J.
Show abstract
Abstract Imaging-based spatially resolved transcriptomics (imSRT) technologies provide high-throughput molecular-resolution spatial characterization of genes within cells. Conventional analysis methods to identify cell-types and states in imSRT data rely on gene count matrices derived from tallying the number of mRNA molecules detected for each gene per segmented cell, thereby overlooking subcellular heterogeneity that can be useful in defining cell states. To take advantage of the molecular-resolution information in imSRT data and potentially identify cell-states based on subcellular heterogeneity, we developed STARIT (Spatial Transcriptomics As Rasterized Image Tensors). STARIT converts transcripts within segmented cells in imSRT data into an image-based tensor representation that can be combined with deep learning computer vision models for downstream analysis. Using simulated and real imSRT data, we demonstrate that STARIT distinguishes transcriptionally distinct cell-types and further separates cell states based on subcellular transcript localization, which conventional gene count analysis fails to capture. By providing a standardized framework to encode subcellular molecular information in imSRT data, STARIT will enable deeper insights into subcellular heterogeneity and enhance the identification and characterization of cell-types and states that are overlooked by gene count representations.
Krieg, R.; Becker, F.; Saenko, S.; Diehl, J.; Stanke, M.
Show abstract
Scaling the structural annotation of protein-coding genes to all eukaryotic genomes remains a major challenge. While recent deep learning methods rival evidence-based pipelines without requiring RNA-seq or alignments, they are entirely supervised. They depend on large, high-quality training sets from diverse genomes, leaving many basal eukaryotic clades without an accurate ab initio gene finder. We present Vipsania, the first unsupervised deep gene finder. A differentiable hidden Markov layer inside a deep sequence model learns to predict gene structures from unannotated genomes alone. Vipsania is pretrained for virtually all eukaryotes and finetunes without supervision on the target genome. It is, on average, more accurate than supervised methods across most clades and avoids the accuracy drop that supervised models suffer on distant target genomes. Vipsania adapts to non-standard genetic codes and provides a fast and highly versatile tool for unbiased, pan-eukaryotic genome annotation. The source code is available at https://github.com/gaius-augustus/vipsania.